Qualcomm AI Engine Direct - Static Decoder Runner Support 16bit KV IO #13127

winskuo-quic · 2025-08-05T09:20:25Z

Summary

Support 16bit KV IO for runner. (Capable to run either 8bit or 16bit)
Adding README for script to run Qwen2.5 0.5B
Improving the PPL score for Qwen2.5 0.5B from 18->12.
Fixing BC CI bug.

Sample Script
python examples/qualcomm/oss_scripts/llama/llama.py -b build-android -s $DEVICE -m SM8750 --prompt "What is 1+1?" --temperature 0 --model_mode kv --max_seq_len 1024 --ptq 16a8w --decoder_model qwen2_5 --eval_perplexity --tasks wikitext --limit 1 --artifact ./16bit_qwen_1024 --enable_masked_softmax --r3

Stats with QNN2.37.0 on SM8750

Accuracy: 12ppl (Align with prepare_pt2e and convert_pt2e)
Token Rate: ~130tok/sec, depending on seq_len.

Test plan

Added E2E test to test_qnn_delegate.py

pytorch-bot · 2025-08-05T09:20:29Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/13127

📄 Preview Python docs built from this PR

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit f7954d3 with merge base c8a0706 ():

NEW FAILURES - The following jobs have failed:

Build documentation / build (buck2) / Build doc (gh)
At least one of the pre-conditions you specified did not hold
pull / unittest / macos / macos-job (gh)
backends/xnnpack/test/ops/test_conv1d.py::TestConv1d::test_qs8_conv1d_batchnorm_seq

This comment was automatically generated by Dr. CI and updates every 15 minutes.

github-actions · 2025-08-05T09:21:03Z

This PR needs a `release notes:` label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

facebook-github-bot · 2025-08-10T04:56:01Z

@cccclai has imported this pull request. If you are a Meta employee, you can view this in D79958806.

cccclai · 2025-08-10T04:57:01Z

There seem to be merge conflict..

facebook-github-bot · 2025-08-11T17:20:07Z

@cccclai has imported this pull request. If you are a Meta employee, you can view this in D79958806.

haowhsu-quic · 2025-08-13T02:07:01Z

Hi @cccclai, I wonder if we could have this merged? It would be great to have this and we can submit PR for statistics real quick.

facebook-github-bot · 2025-08-13T17:00:00Z

@cccclai has imported this pull request. If you are a Meta employee, you can view this in D79958806.

cccclai · 2025-08-13T17:00:10Z

could

Yeah trying to merge but ran into merge conflict, checking again

cccclai

Thanks for making the change!

facebook-github-bot · 2025-08-13T17:35:53Z

@cccclai has imported this pull request. If you are a Meta employee, you can view this in D79958806.

…pytorch#13127) ### Summary - Support 16bit KV IO for runner. (Capable to run either 8bit or 16bit) - Adding README for script to run Qwen2.5 0.5B - Improving the PPL score for Qwen2.5 0.5B from 18->12. - Fixing BC CI bug. Sample Script `python examples/qualcomm/oss_scripts/llama/llama.py -b build-android -s $DEVICE -m SM8750 --prompt "What is 1+1?" --temperature 0 --model_mode kv --max_seq_len 1024 --ptq 16a8w --decoder_model qwen2_5 --eval_perplexity --tasks wikitext --limit 1 --artifact ./16bit_qwen_1024 --enable_masked_softmax --r3` #### Stats with QNN2.37.0 on SM8750 Accuracy: 12ppl (Align with prepare_pt2e and convert_pt2e) Token Rate: ~130tok/sec, depending on seq_len. <img width="1658" height="877" alt="image" src="https://github.com/user-attachments/assets/8fa19068-5613-4329-a527-52f3e02d408f" /> ### Test plan Added E2E test to `test_qnn_delegate.py`

meta-cla bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 5, 2025

winskuo-quic force-pushed the dev1/winskuo/qwen_16_bit branch from 2ee4964 to f9efbc5 Compare August 5, 2025 13:47

winskuo-quic marked this pull request as ready for review August 5, 2025 14:56

winskuo-quic requested a review from cccclai as a code owner August 5, 2025 14:56

winskuo-quic marked this pull request as draft August 5, 2025 16:06

winskuo-quic force-pushed the dev1/winskuo/qwen_16_bit branch from 645a69c to 7bed64b Compare August 6, 2025 05:35

winskuo-quic marked this pull request as ready for review August 6, 2025 07:19

winskuo-quic mentioned this pull request Aug 8, 2025

Qualcomm AI Engine Direct - BC CI Fix and Custom Annotation Fix #13212

Merged

Qualcomm AI Engine Direct - Static Decoder Runner Support 16bit KV IO

f7954d3

winskuo-quic force-pushed the dev1/winskuo/qwen_16_bit branch from 7bed64b to f7954d3 Compare August 11, 2025 03:20

cccclai approved these changes Aug 13, 2025

View reviewed changes

cccclai merged commit d5cd4f3 into pytorch:main Aug 13, 2025
101 of 104 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Qualcomm AI Engine Direct - Static Decoder Runner Support 16bit KV IO #13127

Qualcomm AI Engine Direct - Static Decoder Runner Support 16bit KV IO #13127

Uh oh!

winskuo-quic commented Aug 5, 2025 •

edited

Loading

Uh oh!

pytorch-bot bot commented Aug 5, 2025 •

edited

Loading

Uh oh!

github-actions bot commented Aug 5, 2025

Uh oh!

facebook-github-bot commented Aug 10, 2025

Uh oh!

cccclai commented Aug 10, 2025

Uh oh!

facebook-github-bot commented Aug 11, 2025

Uh oh!

haowhsu-quic commented Aug 13, 2025

Uh oh!

facebook-github-bot commented Aug 13, 2025

Uh oh!

cccclai commented Aug 13, 2025

Uh oh!

cccclai left a comment

Uh oh!

facebook-github-bot commented Aug 13, 2025

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

Qualcomm AI Engine Direct - Static Decoder Runner Support 16bit KV IO #13127

Qualcomm AI Engine Direct - Static Decoder Runner Support 16bit KV IO #13127

Uh oh!

Conversation

winskuo-quic commented Aug 5, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Summary

Stats with QNN2.37.0 on SM8750

Test plan

Uh oh!

pytorch-bot bot commented Aug 5, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/13127

❌ 2 New Failures

Uh oh!

github-actions bot commented Aug 5, 2025

This PR needs a release notes: label

Uh oh!

facebook-github-bot commented Aug 10, 2025

Uh oh!

cccclai commented Aug 10, 2025

Uh oh!

facebook-github-bot commented Aug 11, 2025

Uh oh!

haowhsu-quic commented Aug 13, 2025

Uh oh!

facebook-github-bot commented Aug 13, 2025

Uh oh!

cccclai commented Aug 13, 2025

Uh oh!

cccclai left a comment

Choose a reason for hiding this comment

Uh oh!

facebook-github-bot commented Aug 13, 2025

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

winskuo-quic commented Aug 5, 2025 •

edited

Loading

pytorch-bot bot commented Aug 5, 2025 •

edited

Loading

This PR needs a `release notes:` label